safety concern
Regulating AI 'not the right place to start' says Bailey
Regulating AI'not the right place to start' says Bailey The Governor of the Bank of England has said regulating artificial intelligence (AI) is not the right place to start but instead called first for rigorous testing to find vulnerabilities and safeguards to contain risk. Writing his first-ever article for Substack, Andrew Bailey said the risks around AI were real and increasingly significant. Bailey said the development of AI should not be halted or prohibited - on the contrary, the benefits are immense - but added there must be a system for intervention and to establish boundaries in which AI operates. In recent weeks, the debate about the potential risks surrounding the rapid development of AI and what it means for humanity has intensified. The bosses of leading AI firms such as Anthropic and OpenAI have called for development of the technology to slow down and for an internationally co-ordinated approach assessing risks and putting safeguards in place.
The Download: climate tech companies to watch and AI's discovery problem
The Download: climate tech companies to watch and AI's discovery problem Plus: OpenAI has scrapped a new AI model over safety concerns. With the planet nearing 1.5 C of warming, climate policies being unraveled, and Big Tech backpedaling on its climate ambitions, it can be tempting to give in to climate doom and defeatism. But despite the headwinds, the world has still made incredible progress. That's why publishes its annual list of Climate Tech Companies to Watch . Once a year, we compile a list of 10 companies that we believe have done the most, or have the best chance, to make a real dent in emissions or improve public safety and health. We've now finalized our 2026 list and will publish it on October 6.
As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don't and neither should you Chris Stokel-Walker
As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? The need for independent regulation grows more obvious by the day. We must keep this tech in check before it's too late Fool me more than 16,000 times - as OpenAI agents did to a UN public data hub while repeatedly trying to find its way around the UN's cyber-blocks - and perhaps it's time to admit the system we have for keeping AI agents under control isn't working particularly well. The news about AI systems cropping up in places they shouldn't sounds alarming. Though the description of these as "hacks" is perhaps overstating things, AI has exploited issues in IT systems that humans simply haven't got around to finding.
OpenAI boss and Elon Musk back calls to put brakes on 'reckless' AI development
Growing numbers of US lawmakers are calling for new rules after cases of AI agents going rogue and researchers quitting over safety concerns. Growing numbers of US lawmakers are calling for new rules after cases of AI agents going rogue and researchers quitting over safety concerns. OpenAI boss and Elon Musk back calls to put brakes on'reckless' AI development Sam Altman and Elon Musk have backed a call from the head of Anthropic, Dario Amodei, to "slow the pace" of AI development after he warned that an AI swarm could otherwise become capable of "taking over the entire internet" within a year. In a rare show of unity, the rival technology leaders threw their weight behind the appeal from the founder and chief executive of the artificial intelligence company Anthropic, which came after a series of warnings from AI researchers last week. In an essay, Amodei laid out his thoughts on "why the AI industry should slow down", with a three-part plan for doing so, starting with giving independent monitors constant access to developers' research processes.
The Download: AI's self-improvement problem, and what's driving the heat
Plus: OpenAI has paused some model work over safety concerns. AI's recursive self-improvement might not come so quickly after all The AI industry's boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. But a new study suggests it might take a while to get there. Researchers found that AI agents still can't conduct open-ended AI research--free-form investigations with no clear-cut answers that require the judgment and creativity needed to make genuine breakthroughs. The big question now is how crucial open-ended research is to recursive self-improvement--and whether AI systems can grind their way there without it, simply by improving on narrower tasks. Find out why the results may temper claims that recursive self-improvement is on the horizon .
Amazon's Zoox recalls self-driving vehicles amid emergency response issues
Amazon's Zoox recalls self-driving vehicles amid emergency response issues The Amazon subsidiary company Zoox has said that it will recall its fleet of 105 autonomous vehicles in the United States. The technology company announced the recall on Friday due to mounting concerns that the vehicles may not detect heavy smoke and could impede emergency personnel. Zoox said on Friday that on June 20 an unoccupied Zoox autonomous vehicle encountered heavy smoke that obscured an active emergency fire scene. The Zoox vehicle entered the scene, then braked hard while attempting to steer away, before coming to a stop. The Zoox vehicle, under teleguidance, reversed, after which first responders placed traffic cones at the scene, blocking two of the three lanes.
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
Do LLMs robustly generalize critical safety facts to novel situations? Lacking this ability is dangerous when users ask naive questions--for instance, "I'm considering packing melon balls for my 10-month-old's lunch. What other foods would be good to include?" Before offering food options, the LLM should warn that melon balls pose a choking hazard to toddlers, as documented by the CDC1. Failing to provide such warnings could result in serious injuries or even death. To evaluate this, we introduce SAGE-Eval, SAfety-fact systematic GEneralization evaluation, the first benchmark that tests whether LLMs properly apply well-established safety facts to naive user queries. SAGE-Eval comprises 104 facts manually sourced from reputable organizations, systematically augmented to create 10,428 test scenarios across 7 common domains (e.g., Outdoor Activities, Medicine). We find that the top model, Claude-3.7-sonnet,
Is Your HD Map Constructor Reliable under Sensor Corruptions?
Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is not well understood, raising safety concerns. This work introduces MapBench, the first comprehensive benchmark designed to evaluate the robustness of HD map construction methods against various sensor corruptions. Our benchmark encompasses a total of 29 types of corruptions that occur from cameras and LiDAR sensors. Extensive evaluations across 31 HD map constructors reveal significant performance degradation of existing methods under adverse weather conditions and sensor failures, underscoring critical safety concerns. We identify effective strategies for enhancing robustness, including innovative approaches that leverage multi-modal fusion, advanced data augmentation, and architectural techniques. These insights provide a pathway for developing more reliable HD map construction methods, which are essential for the advancement of autonomous driving technology. The benchmark toolkit and affiliated code and model checkpoints have been made publicly accessible.
Runtime Safety Monitoring of Deep Neural Networks for Perception: A Survey
Schotschneider, Albert, Pavlitska, Svetlana, Zรถllner, J. Marius
Deep neural networks (DNNs) are widely used in perception systems for safety-critical applications, such as autonomous driving and robotics. However, DNNs remain vulnerable to various safety concerns, including generalization errors, out-of-distribution (OOD) inputs, and adversarial attacks, which can lead to hazardous failures. This survey provides a comprehensive overview of runtime safety monitoring approaches, which operate in parallel to DNNs during inference to detect these safety concerns without modifying the DNN itself. We categorize existing methods into three main groups: Monitoring inputs, internal representations, and outputs. We analyze the state-of-the-art for each category, identify strengths and limitations, and map methods to the safety concerns they address. In addition, we highlight open challenges and future research directions.
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
Yueh-Han, Chen, Davidson, Guy, Lake, Brenden M.
Do LLMs robustly generalize critical safety facts to novel situations? Lacking this ability is dangerous when users ask naive questions. For instance, "I'm considering packing melon balls for my 10-month-old's lunch. What other foods would be good to include?" Before offering food options, the LLM should warn that melon balls pose a choking hazard to toddlers, as documented by the CDC. Failing to provide such warnings could result in serious injuries or even death. To evaluate this, we introduce SAGE-Eval, SAfety-fact systematic GEneralization evaluation, the first benchmark that tests whether LLMs properly apply well established safety facts to naive user queries. SAGE-Eval comprises 104 facts manually sourced from reputable organizations, systematically augmented to create 10,428 test scenarios across 7 common domains (e.g., Outdoor Activities, Medicine). We find that the top model, Claude-3.7-sonnet, passes only 58% of all the safety facts tested. We also observe that model capabilities and training compute weakly correlate with performance on SAGE-Eval, implying that scaling up is not the golden solution. Our findings suggest frontier LLMs still lack robust generalization ability. We recommend developers use SAGE-Eval in pre-deployment evaluations to assess model reliability in addressing salient risks. We publicly release SAGE-Eval at https://huggingface.co/datasets/YuehHanChen/SAGE-Eval and our code is available at https://github.com/YuehHanChen/SAGE-Eval/tree/main.